Papers with unsupervised domain adaptation

18 papers
UDAPTER - Efficient Domain Adaptation Using Adapters (2023.eacl-main)

Copied to clipboard

Challenge: Using adapters, unsupervised domain adaptation (UDA) is more parameter efficient and requires large-scale data to be effective.
Approach: They propose to add small bottleneck layers to each layer of a pre-trained language model to make it more parameter efficient by adding adapters.
Outcome: The proposed methods outperform unsupervised domain adaptation methods such as DANN and DSN in natural language inference and sentiment classification tasks.
Unsupervised Domain Adaptation for Question Generation with DomainData Selection and Self-training (2022.findings-naacl)

Copied to clipboard

Challenge: Existing question generation models require large-scale and high-quality training data.
Approach: They propose an unsupervised domain adaptation approach to combat the lack of training data and domain shift issue with domain data selection and self-training.
Outcome: The proposed approach outperforms baselines on three large datasets with different domain similarities, using a transformer-based pre-trained QG model.
Semi-supervised Domain Adaptation for Dependency Parsing (P19-1)

Copied to clipboard

Challenge: Currently, most studies on cross-domain parsing focus on unsupervised domain adaptation . however, unsupervised approaches make limited progress due to the intrinsic difficulty of both domain adaptation and parse.
Approach: They propose a semi-supervised domain adaptation problem for Chinese dependency parsing by using newly-annotated large-scale domain-aware datasets.
Outcome: The proposed method is more effective than direct corpus concatenation and multi-task learning.
Cross-Domain NER using Cross-Domain Language Modeling (P19-1)

Copied to clipboard

Challenge: Existing methods for named entity recognition (NER) use labeled data for both source and target domains.
Approach: They propose to use language modeling as a bridge between NER domains to perform cross-domain and cross-task knowledge transfer.
Outcome: The proposed method extracts domain differences from cross-domain LM contrast, allowing unsupervised domain adaptation while giving state-of-the-art results among supervised domain adapters.
Student Surpasses Teacher: Imitation Attack for Black-Box NLP APIs (2022.coling-1)

Copied to clipboard

Challenge: Existing MLaaS models are vulnerable to imitation attacks, but none of the stolen models can outperform the original black-box APIs.
Approach: They conduct unsupervised domain adaptation and multi-victim ensemble to show attackers could surpass victims.
Outcome: The proposed model outperforms the original black-box models on transferred domains.
Adversarial Domain Adaptation for Machine Reading Comprehension (D19-1)

Copied to clipboard

Challenge: Existing models for machine reading comprehension rely on large amounts of human-annotated in-domain data.
Approach: They propose an unsupervised domain adaptation framework for Machine Reading Comprehension where the source domain has a large amount of labeled data, while only unlabeled passages are available in the target domain.
Outcome: The proposed framework can be generalizable to different MRC models and datasets and can be extended to semi-supervised learning.
Generative Cross-Domain Data Augmentation for Aspect and Opinion Co-Extraction (2022.naacl-main)

Copied to clipboard

Challenge: Existing approaches to perform aspect and opinion co-extraction are difficult due to the lack of fine-grained annotations.
Approach: They propose a framework to transfer knowledge from a labeled source domain to an unlabeled target domain.
Outcome: The proposed framework is more effective than previous domain adaptation methods on three datasets.
Semi-Supervised Domain Adaptation for Emotion-Related Tasks (2023.findings-acl)

Copied to clipboard

Challenge: Semi-supervised domain adaptation (SSDA) is a model trained from a label-rich source domain to a new but related domain with a few labels of target data.
Approach: They propose to decompose the semi-supervised domain adaptation framework into two subcomponents of unsupervised domain adaption (UDA) from the source to the target domain and semi-supervised learning (SSL) in the target.
Outcome: The proposed method is based on the co-learning of multiple classifiers for computer vision tasks and is published in the journal Nature.
Non-Parametric Unsupervised Domain Adaptation for Neural Machine Translation (2021.findings-emnlp)

Copied to clipboard

Challenge: kNN-MT is a non-parametric method that uses nearest neighbor retrieval to translate out-of-domain sentences, rare words, etc.
Approach: They propose a framework that directly uses in-domain monolingual sentences to build an effective datastore for k-nearest-neighbor retrieval.
Outcome: The proposed framework improves translation accuracy with target-side monolingual data while achieving comparable performance with back-translation.
Adversarial and Domain-Aware BERT for Cross-Domain Sentiment Analysis (2020.acl-main)

Copied to clipboard

Challenge: Cross-domain sentiment classification requires large amounts of labeled data.
Approach: They propose to apply a pre-training language model BERT on unsupervised domain adaptation . they propose to distill domain-specific features in a self-supervised way .
Outcome: The proposed model outperforms state-of-the-art methods on Amazon dataset . it can be applied to the unsupervised domain adaptation task without domain awareness .
DUQGen: Effective Unsupervised Domain Adaptation of Neural Rankers by Diversifying Synthetic Query Generation (2024.naacl-long)

Copied to clipboard

Challenge: State-of-the-art rankers pre-trained on large task-specific training data such as MS-MARCO exhibit strong performance on various ranking tasks without domain adaptation, also called zero-shot.
Approach: They propose a method to generate unsupervised domain adaptation for ranking using large-scale task-specific training data such as MS-MARCO and Wikipedia retrieval.
Outcome: The proposed method outperforms all zero-shot baselines and significantly outperfies the SOTA baselines on 16 out of 18 datasets, for an average of 4% relative improvement across all datasets.
Unsupervised Domain Adaptation for Text Classification via Meta Self-Paced Learning (2022.coling-1)

Copied to clipboard

Challenge: Recent methods addressing unsupervised domain adaptation for textual tasks extracted domain-invariant representations through balancing between multiple objectives to align feature spaces between source and target domains.
Approach: They propose to use meta-learning framework to train a neural network-based self-paced learning procedure in an end-to-end manner.
Outcome: The proposed method significantly improves performance on target domains, surpassing state-of-the-art approaches.
Transferable End-to-End Aspect-based Sentiment Analysis with Selective Adversarial Learning (D19-1)

Copied to clipboard

Challenge: Existing methods to extract aspects and sentiments are limited due to lack of annotated sequence data.
Approach: They propose a Selective Adversarial Learning method to align latent correlation vectors . they propose tagging a set of aspect boundary tags and sentiment tags to create a joint label space .
Outcome: The proposed method can learn weights for words to achieve fine-grained adaptation.
Multi-Source Domain Adaptation with Mixture of Experts (D18-1)

Copied to clipboard

Challenge: Existing methods for domain adaptation from multiple sources are designed to transfer supervision from a single source domain.
Approach: They propose to capture the relationship between a target example and different source domains by a point-to-set metric.
Outcome: The proposed method outperforms baselines and can handle negative transfer.
Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval (2021.emnlp-main)

Copied to clipboard

Challenge: Using self-training to train unsupervised domains can be expensive, resulting in poor generalization due to distributional shift.
Approach: They propose to use unaligned data to train unsupervised domain adaptation models using cheap synthetically generated labeled data.
Outcome: The proposed method significantly outperforms self-training on question generation and passage retrieval domains and on MLQuestions and PubMedQA.
Task Refinement Learning for Improved Accuracy and Stability of Unsupervised Domain Adaptation (P19-1)

Copied to clipboard

Challenge: Existing approaches to domain adaptation (DA) require labeled data that can be found in only a handful of domains.
Approach: They propose a task-refinement learning approach to solve pivot detection problems . they propose to train PBLM models with gradually increasing information exposed about each pivot .
Outcome: The proposed approach achieves state-of-the-art accuracy in six domain adaptation setups for sentiment classification.
Feature Adaptation of Pre-Trained Language Models across Languages and Domains with Robust Self-Training (2020.emnlp-main)

Copied to clipboard

Challenge: Adapting pre-trained language models (PrLMs) to new domains has gained much attention . Adaptation of PrLMs to newdomains is important, but requires fine-tuning .
Approach: They propose to use PrLMs to adapt to new domains without fine-tuning . they use class-aware feature self-distillation to learn discriminative features .
Outcome: The proposed model can learn discriminative features from pre-trained language models without fine-tuning.
Iterative Nearest Neighbour Machine Translation for Unsupervised Domain Adaptation (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for supervised domain adaptation of machine translation focus on fine-tuning, which is non-extensible.
Approach: They propose to perform unsupervised domain adaptation in a non-parametric manner by using in-domain monolingual data and performing nearest neighbour inference on both forward and backward directions.
Outcome: The proposed method significantly improves the in-domain translation performance and achieves state-of-the-art results among non-parametric methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations